Conversation
Bug (fixed): for unpaired data where the same clusters appear in both groups, the permutation test reshuffled labels within each cluster. That tests the strong null of no effect in any participant, so participant-specific effects that average to zero still caused rejections: 20% at the 5% level. It now swaps each cluster's whole control and test sets, the direct analogue of the paired sign flip. After the fix it rejects 3% at the 5% level and 10% at the 10% level. See dabest/_effsize_objects.py:1885.
Baseline error curve documentation did not mention the details that make its use questionable to paired or clustered analyses. show_baseline_ec docstring (dabest/_effsize_objects.py:1584): now explains what the curve actually computes, states plainly that it's always an unpaired self-comparison regardless of paired, and calls out the scale mismatch with cluster_col. cluster_col docstring (dabest/_api.py:92): adds a pointer to that same caveat, since a user reading about cluster_col alone wouldn't otherwise know it affects the baseline curve disproportionately. Plot Aesthetics tutorial (08-plot_aesthetics.ipynb): rewrote the "Baseline error curve" section and added a new subsection with a worked, executed example. It builds a 15-participant, 3-pairs-each dataset and plots the same comparison with id_col alone versus id_col plus cluster_col, side by side. The real paired contrast stays essentially unchanged between the two; the baseline curve at the "Control" position visibly widens, from a 95% width of about 1.2 to about 2.0 in this example. That's the same phenomenon you flagged, reproduced deliberately and small enough to read at a glance, rather than buried in a large real dataset. Changelog: added a Documentation entry alongside the existing cluster-aware bootstrap entry.
Added new tutorial to demonstrate the use and value of the cluster-aware bootstrap feature, and added hooks in associated notebooks.
Ensure plot labels reflect both observations and cluster sample sizes.
…ter sample sizes Hesterberg expanded percentile adjustment to expand coverage for small numbers of clusters. Active when cluster_col is assigned, opt-out argument also available.
Degrees of freedom adjustment for percentile expansion switched from Hsu to Welch-Satterthwaite to avoid over-conservatism. Plots also now indicate expanded intervals with thinner lines, also editable.
|
Check out this pull request on See visual diffs & provide feedback on Jupyter Notebooks. Powered by ReviewNB |
Extra gap added to accommodate cluster sample size in plots needed to be rationalised.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add cluster-aware bootstrapping and permutation tests (
cluster_col)Addresses enhancement feature request issue #241.
Summary
DABEST's bootstrap and permutation test treat every row of
id_col(or, for unpaired data, every observation) as an independent sampling unit. Some repeated-measures designs break that assumption: when a participant contributes several paired sets, for example several stimuli rated under each condition. Analysing these designs using DABEST represents pseudoreplication, with the impact that confidence intervals are estimated as too narrow.This feature adds a
cluster_colargument todabest.load()naming the real independent sampling unit. When it is set, whole clusters are resampled and reshuffled, and intervals are expanded for the number of clusters, so that intervals reflect between-cluster correlation and keep close to nominal coverage. When it is not set, nothing changes.What changed
dabest.load(..., cluster_col=None, cluster_ci_expansion=True): works with unpaired and paired data (id_colidentifies the pairs,cluster_colthe units they are nested in), shared-control and multi-groupidx, proportional (binary) data, delta-delta and mini-meta.ci) is unchanged; the level actually read is reported asci_expanded, with the unexpanded limits alongside.cluster_ci_expansion=Falseturns it off.n_clusters,ci_expandedand unexpanded-limit columns (bca_low_unexpanded,pct_high_unexpanded,bec_bca_low_unexpanded, ...) in.results, which appear only when they apply;ci_expandedin.statistical_tests;TwoGroupsEffectSizeandPermutationTestacceptcontrol_clusters/test_clusters(andTwoGroupsEffectSizeacceptscluster_ci_expansion).cluster_colset, each group label reads(N=<observations>,/n=<clusters>)on two lines. In vertical Cumming plots (including two-column Sankey plots), the gap between the raw-data and contrast axes is sized to the labels.contrast_expanded_errorbar_kwargsin.plot()andexpanded_errorbar_kwargsinforest_plot().x,yor anidxgroup, and pairs whose rows carry different cluster labels; a warning when too few clusters remain for any bootstrapinterval to be reliable.
Point estimates are unaffected. The parametric and rank-based tests in
.statistical_testsare unchanged and still treat observations as independent.Accuracy checks
Documentation
cluster_colwalkthrough.show_baseline_ecdocstring: the baseline error curve is always an unpaired self-comparison, so withcluster_colit can be much wider than the real paired contrasts.load(),plot(),forest_plot(),show_sample_size,TwoGroupsEffectSizeandPermutationTest, plus a CHANGELOG entry.Limitations
Testing
nbs/tests/test_cluster_bootstrap.py: 34 tests. They cover:Every test, API and tutorial notebook executes under
nbdev-test. All notebooks are clean and in sync with the library.Backward compatibility
cluster_coldefaults toNone. Every new code path is gated on it being set, including the expansion, the two-part interval, the plot-label and layout changes, and the new result columns.